# UNITE Corpus
UNiversally Inclusive Technologies to practice English


## Corpus summary

Language: English (learner English)  
Text producers: Italian university students (age 19–25)  
Text type: conversations between humans and AI-based chatbots
Tasks performed by learners: small talk and role play
Chatbots: ChatGPT, Pi.AI
Size: 326 texts; 722,537 tokens 
Collection period: May–December 2024  
Institutions: University of Bologna, University of Macerata, University of Naples “L’Orientale”  
Annotation: **metadata**   + - **Semantic annotation (DIS-TAG)**

Full documentation of the annotation scheme is available in the associated Zenodo record:
	DIS-TAG: De Brasi, V., & Mongibello, A. (2026). Normative DIS-TAG_UNITE [Data set]. Zenodo. https://doi.org/10.5281/zenodo.18960403
	
## Project Information

Project title: UNITE – UNiversally Inclusive Technologies to practice English  
Funding: European Union – NextGenerationEU, Ministero dell'università e della ricerca (MUR) 
Programme: Italian National Recovery and Resilience Plan (PNRR), PRIN 2022  
Project code: 2022JB5KAL  
CUP: J53D23008070006
Project website: https://site.unibo.it/unite/en

## Corpus Composition

The corpus consists of a total of **329 annotated conversations** collected from three Italian universities:

- **University of Naples “L’Orientale” (unior)**: 72 conversations  
- **University of Bologna (unibo)**: 166 conversations  
- **University of Macerata (unimc)**: 91 conversations  

Each file represents a **single interaction** between a learner and an AI chatbot.

## File Naming Convention

Files are named according to the following structure:

- `unior_xx` → University of Naples “L’Orientale”  
- `unibo_xx` → University of Bologna  
- `unimc_xx` → University of Macerata  

Where `xx` indicates the conversation number.

Examples:

- `unior_1.txt`
- `unibo_25.txt`
- `unimc_10.txt`

This naming convention allows users to identify the institutional source of each interaction.

## File Structure

annotated_UNITE_corpus/

unior_1.txt  
unior_2.txt  
...  
unibo_1.txt  
...  
unimc_1.txt  
...

## Chatbot Systems

The dataset includes interactions with the following AI-based conversational agents:

- ChatGPT  
- Pi.AI  

These systems were used by learners for English language practice tasks.

## Ethical Considerations

All data have been anonymized.  
The corpus does not contain personally identifiable information.

The dataset is shared for research and educational purposes.

The consent forms and student instructions are available on the project website: https://site.unibo.it/unite/en/join-us

## If you use this source please cite

De Brasi, V., Cecchini, S., & Mongibello, A. (2026). annotated_UNITE_corpus [Data set]. Zenodo. https://doi.org/10.5281/zenodo.19063359


These Guidelines were developed within the framework of the PRIN 2022 project UNITE – Universally
Inclusive Technologies to Practice English, funded by the Italian Ministry of University and Research under
the European Union – NextGenerationEU program. Project Code 2022JB5KAL; CUP J53D23008070006.